Statistics in Medicine
○ Wiley
Preprints posted in the last 30 days, ranked by how well they match Statistics in Medicine's content profile, based on 40 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Mell, L. K.
Show abstract
In competing risks settings, covariate effects and group comparisons are usually assessed one event at a time - through log-rank or Cox tests on the cause-specific hazards, or Gray's test or Fine-Gray regression on a cumulative incidence function (CIF). This can obscure a clinically important quantity: the ratio between the event of interest and the competing event, since groups may differ little on the individual events yet differ sharply in their ratio. The generalized competing event (GCE) framework makes this ratio the object of inference; on the cause-specific scale the hazard ratio omega+(t) = lambda_1(t)/lambda_2(t) is estimated efficiently from a single stacked (Lunn-McNeil) model. We extend the framework to two scales that describe realized incidence. The subdistribution hazard ratio omega-tilde+(t) = lambda-tilde_1(t)/lambda-tilde_2(t) is estimated by a stacked, risk-set-weighted extension of the Lunn-McNeil construction; the cumulative-incidence ratio rho(t) = F_1(t)/F_2(t) - the odds that a subject's realized event by time t is the event of interest - by jackknife pseudo-observation regression of the Aalen-Johansen estimator. We relate the three contrasts: rho equals omega+ exactly under proportional cause-specific hazards, and equals omega-tilde+ only in the small-time limit under proportional subdistribution hazards, drifting toward 1 thereafter. The orthogonality that makes omega+ efficient is lost on both cumulative-incidence scales - omega tilde+ through overlapping weighted risk sets and shared censoring weights, rho through the shared all-cause survivor - so each carries a covariance term that must be handled and that bounds efficiency relative to the hazard-scale test. We derive the corresponding variances, study operating characteristics by simulation, illustrate on hypothetical prostate and head-and-neck cohorts, and provide an implementation in the gcemod R package.
Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.
Show abstract
Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([≥] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([≥] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.
Chen, T.; Voorhies, K.; Reeson, A.; Seo, S.; Lee, S.; Hahn, G.; Hecker, J.; Prokopenko, D.; Hoth, K.; Kelly, R.; Lasky-Su, J. A.; Weiss, S.; Lange, C.; Lutz, S.
Show abstract
Mendelian Randomization (MR) is a popular tool for inferring causal relationships between traits using genetic variants as instrumental variables. These methods have been extended to also determine the direction of causality. However, causal direction cannot be inferred from a statistical test or estimation procedure (i.e. from data alone) without further assumptions and the methods operating characteristics and relative performances are not well understood. We conducted a comprehensive simulation study to illustrate this issue by evaluating type I error and power of 17 summary-based MR methods for inferring the effect direction. These methods fall within three methodological families: MR Steiger, Causal Direction (CD), and bidirectional MR approaches, with scenarios ranging across combinations of horizontal pleiotropy, unmeasured confounding, measurement error, longitudinal feedback, and varying sample sizes. While most methods achieved sufficient power levels under the alternative hypothesis in most scenarios, we found that every method was susceptible to inferring the wrong causal direction or under powered, and no method consistently maintained both correct type 1 error control and high power. In our applications, we evaluated the effect direction between the trait pairs body mass index (BMI) and major depressive disorder (MDD) and between BMI and asthma. To help researchers to evaluate the 17 methods to infer the effect direction and consider these challenges in their own data, we have developed MRdirection, an R package that runs the simulation studies examining the 17 directional MR methods across different user-defined scenarios. Our study, together with the accompanying R package, provides researchers with a tool for examining directional MR methods given different underlying assumptions.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Pocuca, T.; Pare, G.; Bolker, B. M.
Show abstract
Accurate normalization is essential for differential expression analysis of RNA-sequencing data. Popular normalization methods such as the median-of-ratios and trimmed mean of M-values do not leverage information from the experimental design. This may be inefficient in experiments with large-scale systematic expression changes or complex designs. Here, we introduce design-informed size factor estimation (disize), a normalization method that uses information from the experimental design to improve accuracy. disize uses a modified generalized linear mixed model to robustly distinguish between biological signal and sample-specific size factors. We also propose a mechanistically justified data-generating process for RNA-sequencing counts that is derived from previous models of transcription and sequencing. Through simulations based on this data-generating process and validating on true RNA-seq data, we show that disize recovers size factors more accurately than existing methods, particularly in challenging scenarios with low gene expression and a high proportion of differentially expressed genes; this in turn improves downstream analysis. disize provides a robust and accurate approach to normalization, highlighting the significant benefits of integrating experimental design information directly into normalization for transcriptomic datasets. Author summaryIn transcriptomic analysis, normalization adjusts for technical biases arising from library preparation and sequencing. Methods implemented in widely used packages like DESeq2 and edgeR ignore information in the experimental design during normalization. Incorporating information from the experimental design into a normalization method has the potential to yield more accurate results. To do this, we developed a new method, design-informed size factor estimation (disize), that uses a statistical model to jointly account for the biological signal defined by the design and the sample-specific batch effect. By separating the biological variation into its components, disize can more robustly estimate the batch effect. To validate our approach, we constructed a flexible simulation framework relying on a mechanistically justified data-generating process for RNA-seq data. Our benchmarks on both simulated and true RNA-seq data show that disize recovers the true size factors more accurately than existing methods, particularly in challenging scenarios with low counts or a high proportion of differentially expressed genes. This improved normalization yields more reliable downstream results in differential expression analysis.
Mason, A. C.; Ballabio, G.; Paz, V.; Sofat, R.; Garfield, V.
Show abstract
Mendelian randomization (MR) is widely used to infer causal relationships using genetic variants as instrumental variables, yet the selection of genetic instruments is not always given sufficient attention. Many MR studies rely on default linkage disequilibrium (LD) clumping parameters (r2 <0.001, 10,000 kb), as implemented in commonly used tools, without assessment of their suitability for specific exposures. We investigated whether this approach yields optimal instruments or whether a more pragmatic strategy yields stronger instruments. Using UK Biobank data, we examined three distinct exposure types-circulating amino acids, body mass index (BMI), and major depressive disorder (MDD). For each phenotype, we systematically varied LD clumping thresholds (r2 and genomic distance) and evaluated each instrument via both their average strength (F-statistic) and total strength (R2). Across all phenotypes, optimal instruments differed from default parameters and varied by exposure. For amino acids and BMI, more stringent LD thresholds (r2=0.00001) combined with larger clumping windows improved instrument strength, whereas for MDD, a highly polygenic, binary trait, smaller windows with stringent r2 maximized variance explained while maintaining F-statistics above the desired threshold (>10). Notably, increasing the number of SNPs did not consistently improve instrument quality, highlighting a trade-off between instrument strength and potential pleiotropy. We demonstrate that universal reliance on default LD clumping parameters can lead to suboptimal instruments. We propose a pragmatic framework for instrument selection based on empirical evaluation of strength metrics, improving the robustness and transparency of MR analyses across different exposure types.
Ohno, K.
Show abstract
Background. Sequential monitoring of risk-adjusted mortality in intensive care typically relies on alarm-generating control charts - the risk-adjusted CUSUM, EWMA, or VLAD. These charts signal deterioration but do not return what clinicians and registry stewards ultimately need to interpret: a calibrated, continuously updated estimate of the standardized mortality ratio (SMR) itself. Methods. STREAM-SMR is a conjugate gamma-Poisson dynamic generalized linear model in which the latent log-SMR evolves through a discount factor delta and the alarm statistic is the posterior exceedance probability P(SMR > 1). The construction is closed-form, exact for zero-death months, and computationally trivial at registry scale. Under a protocol frozen before any evaluation runs and calibrated to published national ICU registry aggregates, we compared STREAM-SMR (delta in {0.90, 0.95, 0.97}) with the risk-adjusted CUSUM and risk-adjusted EWMA across three facility-volume strata (50, 200, and 800 annual admissions) and five change scenarios (sustained steps, gradual drift, transient deterioration, and improvement). All methods were Monte-Carlo-calibrated to a common 5% false-alarm probability over a 60-month horizon, with 1,000 replications per cell. Results. At delta = 0.90, STREAM-SMR matched the detection frontier of the risk-adjusted CUSUM to within one to two months across sustained-shift scenarios - median delay for an SMR step to 1.5 of 6 versus 5 months in large facilities, 14 versus 14 in medium, and 20 versus 22 in small - while returning filtered SMR estimates whose 95% credible intervals held at least 91% empirical coverage in every scenario-stratum cell. Empirical false-alarm probabilities were close to the 5% nominal target for all methods (range 0.035-0.066). For a three-month transient deterioration, detection was faster with STREAM-SMR conditional on occurring, but overall detection probability favored the CUSUM in large facilities. Conclusions. STREAM-SMR unifies monitoring and estimation in a single Bayesian object: for a detection-delay premium of at most one to two months against the theoretically optimal CUSUM, it returns an interpretable, uncertainty-quantified SMR trajectory at every time point. The discount factor is an explicit dial between estimate smoothness and detection speed. Simulation code and the frozen protocol are publicly archived.
Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.
Show abstract
Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.
Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.
Show abstract
Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.
Cai, H.; Yu, W.; Lu, R.; Chattopadhyay, I.; Zhang, X.; Liu, J.
Show abstract
Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existing generative methods, particularly generative adversarial network (GAN)-based approaches, can model complex distributions, but their estimated dependence structures are often difficult to interpret and their performance may be unstable or prone to overfitting in modestly sized datasets. Here we show that eCDF-copula, a statistically rooted approach using the empirical cumulative distribution function (eCDF) and copula modeling, preserves within- and between-visit dependence structure. To handle pervasive missing data, we propose a two-stage strategy combining multiple imputation with copula-based synthesis, enabling a variance decomposition that quantifies replication variability across methods. We benchmarked the proposed approach against four established methods on two longitudinal clinical datasets spanning markedly different sample sizes (n = 120 vs. n = 3, 612). eCDF-copula achieved resemblance and utility exceeding those of state-of-the-art synthetic data methods, while maintaining comparable privacy.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Isaiev, B.; Stukalova, I.
Show abstract
Background: The growing burden of lifestyle-related chronic diseases has increased the need for clinically interpretable decision-support tools capable of integrating artificial intelligence with evidence-based preventive nutrition. Although machine learning has shown considerable potential for health risk prediction, most existing approaches remain limited to isolated predictive models or conventional nutritional software, with little integration of multidimensional clinical assessment and personalized recommendations. Objective: To develop and internally validate NutrIA, a hybrid web-based Clinical Decision Support System (CDSS) that combines machine learning, validated clinical assessment, structured clinical reasoning and personalized nutritional recommendations for preventive medicine. Methods: NutrIA was developed using harmonized data from the National Health and Nutrition Examination Survey (NHANES, 1988 to 2018). A supervised machine learning model was trained to estimate 5-, 10- and 20-year all-cause mortality risk and subsequently integrated with an adaptive clinical questionnaire, validated screening instruments, nutritional indicators, dietary clustering, clinical phenotyping and a transparent rule-based recommendation engine within a unified web-based platform. Results: The predictive model achieved ROC-AUC values of 0.894, 0.914 and 0.923 for 5-, 10- and 20-year mortality prediction, respectively. The implemented CDSS incorporates an adaptive questionnaire (151 items), 39 validated clinical assessment instruments, 17 clinical phenotypes and 31 dietary clustering modules to generate individualized nutritional and lifestyle recommendations together with an automated clinical report. The integrated framework translates probabilistic risk estimates into clinically interpretable decision support for personalized preventive nutrition. Conclusions: NutrIA demonstrates the technical feasibility of integrating machine learning with knowledge-based clinical reasoning within a single web-based CDSS for preventive nutrition. Although external validation and prospective clinical evaluation are required before routine implementation, the proposed architecture represents a promising step toward clinically interpretable artificial intelligence for personalized nutritional care.
Robinson, C. L.; Turkington, D.; Lee, L.; Kritzman, M.; Yong, R. J.
Show abstract
Accurate prediction of individual medical outcomes is essential for optimizing treatment allocation amid rising costs, coverage denials, and limited clinical resources. Traditional predictive models, including regression and neural networks, rely on average effects and cannot tailor predictions to the specific circumstances of individual cases. We present relevance-based prediction (RBP), a model-free method that predicts outcomes as weighted averages of observed cases, with weights determined by a rigorously defined measure of relevance. Unlike model-based methods that rely on fixed calibrated parameters, RBP revisits the original data for each prediction and customizes both the cases and variables used. Applied to opioid treatment, RBP provides case-specific insights unavailable from conventional models, including how each prior case informs a prediction, how each variable affects its reliability and value, and how reliable the prediction is before it is made. These individualized insights may prevent misleading average-based decisions and reduce harmful or suboptimal treatment.
Chervet, S.; Layan, M.; Boëlle, P.-Y.; Guedj, J.; van der Werf, S.; Kerneis, S.; Sermet-Gaudelus, I.; Cauchemez, S.; Opatowski, L.
Show abstract
Longitudinal household studies, combined with mathematical modeling, are widely used to characterize the drivers of respiratory pathogen transmission, including the effects of age and symptoms. In practice, household recruitment protocols vary across studies, potentially introducing biases into observed data. However, these biases are typically overlooked in statistical inference, and their impact on parameter estimates remains unknown. Here, we use synthetic household outbreak data simulated under different recruitment protocols to evaluate how recruiting through infected children affects estimates of age-specific infectiousness and susceptibility. We show that, under child-based recruitment, the standard likelihood, which accounts only for transmission dynamics, leads to underestimating child infectiousness and overestimating child susceptibility by more than 30%. We then propose a novel estimation framework that explicitly incorporates the household recruitment process into the likelihood and show that it substantially reduces these biases. Applying this new approach to a French household study conducted during the COVID-19 pandemic, we estimated that children under 6 had 49% lower infectiousness than teenagers and adults during the Alpha wave, whereas no difference was observed during the Omicron wave. This study demonstrates that ignoring recruitment protocols can bias key epidemiological parameter estimates and highlights the importance of accounting for study design.
Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.
Show abstract
Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.
Weidemueller, P. H.; Esquivel Gomez, L. R.; Rodriguez-Barraquer, I.; Mueller, N. F.
Show abstract
Tracking how an infectious disease spreads in time and space relies on several distinct sources of surveillance data, reported case counts, viral concentrations in wastewater, seroprevalence surveys, and pathogen genomic sequences, each of which is imperfect and captures only part of the underlying transmission process. These data streams are typically analyzed separately or with highly parameterized, disease-specific models, making it difficult to combine their complementary strengths. Here we present MASCOT-DataStreams (MASCOT-DS), a BEAST2 software package that extends the structured coalescent model MASCOT to jointly infer prevalence over time and transmission rates between locations from any combination of case counts, wastewater concentrations, seroprevalence surveys, and pathogen phylogenies. Using simulated outbreaks in structured populations, we show that MASCOT-DS accurately recovers true prevalence trajectories and between-location migration rates. We then apply MASCOT-DS to genomic, case count, wastewater, and seroprevalence data from the SARS-CoV-2 Epsilon wave (winter 2020-21) in three San Francisco Bay Area counties, reconstructing county-level prevalence dynamics and quantifying transmission within and into the region. By systematically removing individual data streams, we find that genomic data are uniquely required to estimate transmission between locations, while seroprevalence data are essential for anchoring the overall magnitude of an outbreak; case counts and wastewater concentrations play largely interchangeable roles in capturing outbreak shape. These results demonstrate that integrating complementary epidemiological data streams substantially increases the certainty of transmission dynamics estimates compared to relying on any single data stream, and provides a framework for evaluating the added value of different surveillance strategies.
Merlo, J.; Bashir, N. Z.; Rodriguez-Lopez, M.; Khalaf, K.; Öberg, J.; Perez-Vicente, R.
Show abstract
Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) describes health inequalities through three components: (i) specific contextual effects (SCE), (ii) general contextual effects (GCE), and (iii) discriminatory accuracy of the context. We present Simple-Means MAIHDA (S-MAIHDA), which estimates each stratum directly from its observed individuals, with no distributional assumption. The observed proportions are unbiased whatever the stratum size, and their confidence intervals report the uncertainty honestly. S-MAIHDA operationalises the three components on the probability scale. The SCE are the raw and standardised stratum prevalences and the modification of the sociodemographic average differences by the area. The GCE are the variance partition coefficient (VPC) and the contextual structuring of the between-stratum inequality, expressed as the contextual clustering of inequalities, the additive sociodemographic differences, and the contextual modification of inequalities (CMI). The contextual discriminatory accuracy is expressed by the area under the ROC curve (AUC), and the sensitivity and specificity at the population prevalence as the threshold for a possible intervention. Because its estimates are the observed data themselves, S-MAIHDA is the canonical description, and the compare diagnostic quantifies how Random-Effects MAIHDA (RE-MAIHDA), the usual implementation, departs from it: RE shrinkage pulls small strata towards the overall mean and can hide the very inequalities the analysis seeks. The approach is implemented in the smaihda Stata command and reproduced in free Python code. We illustrate S-MAIHDA on register data from Malmo, Sweden (43,291 individuals; 300 area-sociodemographic strata), showing how the three components separate two contrasting outcomes: psychotropic medication use, almost purely sociodemographic, stable across areas, with weak contextual structuring (VPC {approx} 4%, CMI {approx} 0%); and choice of a private general practitioner, strongly geographical (VPC {approx} 11%, CMI {approx} 17%), with the sociodemographic differences reshaped and amplified in wealthy areas. RE-MAIHDA attenuated inequalities. For describing inequalities, S-MAIHDA preserves what the data show.
Scherer, L. D.; Matlock, D. D.; Cronin, J.; Gritz, M.
Show abstract
Multi-Cancer Detection (MCD) tests can detect more than 50 different types of cancer using a blood test. Recently passed law in the U.S. guarantees that Medicare will pay for these tests when they are FDA approved and show evidence for clinical benefit. This manuscript provides estimates of the cost of MCD tests to Medicare under different assumptions of cost per test, eligibility, and screening uptake in the eligible population. This manuscript additionally estimates the cost of follow-up testing resulting from false positive results, which are considered avoidable costs caused by the screening test.
Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.
Show abstract
Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.
Jhand, A. S.; Greenwald, M. K.
Show abstract
Quantifying decision-making in experimental settings that mimic real-world conditions may provide insights into mechanisms underlying addiction. This study developed a computational model of opioid-seeking behavior. Out-of-treatment persons who regularly used heroin were stabilized on buprenorphine 8mg/day to minimize opioid withdrawal. Across programmatically-linked studies, three experimental conditions presented differing money vs. opioid unit amounts that could be earned per trial ($2 vs. 1-mg hydromorphone, n=23; $2 vs. 2-mg hydromorphone, n=36; $4 vs. 2-mg hydromorphone, n=24), controlling other factors. Progressive ratio schedules on each choice option required increasing effort across trials to earn the same amount. Trial-level outcomes were decision latency and choice on each option, and session-level outcomes were drug-money latency and breakpoint difference scores. A Markov computational model was used to predict the probability of choosing the same option as the previous trial (vs. switching). Model inputs included effort discrepancy (between earning the same vs. other commodity on next choice) and logarithm of the ratio of decisional speed (current vs. previous choice). Participants who more rapidly chose hydromorphone vs. money made more consecutive drug choices and expended greater effort earning hydromorphone. First-trial hydromorphone choice predicted continued effortful opioid-seeking. Participants repeated choices on 80% of trials; the model accurately predicted stick vs. switch behavior on 93% of trials. Participants typically repeated choices when faced with lower effort discrepancies and higher hydromorphone dose (2-mg vs. 1-mg). In conclusion, a Markov computational model accurately predicted effortful behavior in a choice paradigm that mimics real-world decisions between opioid and nondrug reinforcers.